The Annals of Applied Statistics
● Institute of Mathematical Statistics
Preprints posted in the last 7 days, ranked by how well they match The Annals of Applied Statistics's content profile, based on 19 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Pizarro Galleguillos, F.; Bhonsale, S.; VAN IMPE, J.
Show abstract
The dynamics of gene regulatory networks are governed by intrinsic noise, stemming from the random nature of biochemical reactions, and by extrinsic noise, arising from fluctuations in cellular components and environmental conditions. Together, these sources can compromise the reliability of predictive computational models if not properly accounted for, and capturing both effects within a single framework remains a non-trivial task in computational biology. In this work, we propose an uncertainty quantification framework that addresses these two contributions jointly: intrinsic stochasticity is described through a partial integro-differential equation (PIDE) for the protein probability density function, whereas extrinsic noise is represented as parametric uncertainty in the kinetic parameters. The propagation of the uncertainty is carried out via an intrusive polynomial chaos expansion (PCE), in which the PCE coefficients are obtained from a stochastic Galerkin projection of the PIDE, yielding a coupled deterministic system that is solved with standard numerical methods. We illustrate the approach on a positive autoregulatory gene network with one and two uncertain kinetic parameters. The proposed approach accurately reproduces the mean, variance, and full protein probability density function, including the bimodal distributions, at a substantially lower computational cost.
Li, X.; Wei, P.
Show abstract
Causal mediation analysis is widely used to identify biological pathways linking exposures to outcomes, but most methods assume homogeneous mediation effects across individuals. In high-dimensional omics settings, this assumption can mask important heterogeneity driven by demographic, genetic, or environmental factors. We propose the M-high-learner, a flexible framework for detecting heterogeneous mediation effects with high-dimensional mediators. The method identifies mediators with subgroup-specific indirect effects while distinguishing them from null or homogeneous signals and controlling the type I error rate. It is computationally efficient, scalable, and yields interpretable sub-types. Simulation studies show that the proposed approach achieves high power while maintaining accurate error control. Applications to the Framingham Heart Study and the Multi-Ethnic Study of Atherosclerosis reveal that the mediation role of gene expression in sexs effect on high-density lipoprotein varies across subgroups defined by body mass index and age. Our framework provides a practical tool for uncovering heterogeneous biological mechanisms in high-dimensional genomic studies. Author SummaryBiological processes linking risk factors to disease often differ across individuals, but many existing methods assume these processes are the same for everyone. This can hide important differences between groups. We developed a powerful method to identify when these pathways vary across subgroups using large-scale molecular data. Our approach detects differences in how intermediate biological factors contribute to outcomes in populations defined by characteristics such as age and body mass index. Applying our method to population studies, we found that some biological pathways operate differently across groups, suggesting that key mechanisms may be missed when differences are ignored. Our work provides a tool to better understand how disease-related processes vary across individuals, which may support more targeted and personalized approaches to health research.
Choi, L.; McNeer, E.; Beck, C. A.; Neul, J. L.
Show abstract
Bayesian borrowing of external information can improve trial efficiency, particularly in pediatric and rare disease settings where patient populations are limited, but may introduce bias and inflate the Type~I error rate when the trial differs from external studies. Recent U.S. Food and Drug Administration (FDA) draft Bayesian guidance emphasizes careful evaluation of external information, prior specification, and assessment of operating characteristics. This paper compares three meta-analytic-predictive (MAP)-based methods for Bayesian borrowing: the MAP prior, robust MAP (RMAP) prior, and self-adapting mixture (SAM) prior. An adaptive platform trial design in Rett syndrome is used as a case study. Simulation studies evaluate frequentist operating characteristics under varying prior--data conflict, between-study heterogeneity, treatment effects, and clinically significant differences (CSDs) for the SAM prior. The MAP prior achieved the greatest efficiency when external and current data were compatible but exhibited the largest bias under substantial prior--data conflict. The RMAP priors improved robustness through fixed robust-component weights, whereas the SAM prior adaptively adjusted borrowing and was less sensitive to prior--data conflict while retaining efficiency gains when the data were compatible. Although the CSD influenced the degree of adaptive borrowing, as reflected by effective sample size, it had only a modest impact on frequentist operating characteristics. Sensitivity analyses using a skeptical robust component yielded similar qualitative conclusions, while accentuating the differences between the MAP and RMAP priors. These findings provide guidance for evaluating and selecting MAP-based borrowing strategies before trial implementation, particularly in rare disease settings, consistent with current FDA recommendations.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Alve, S. R.; Rahman, S.; Meem, S. M. A. C.
Show abstract
A dental AI system and a dentist reading the same radiographs form a paired comparison. Published comparative studies often report the two arms separately against a reference standard, leaving the joint pattern of correctness between them unavailable for secondary paired inference. We show what that omission costs. The accuracy difference remains exactly identified; its sampling variance does not, so the report contains the estimate and not its uncertainty. On a study of 282 units, two published accuracies are consistent with 38 distinct joint tables whose confidence intervals differ in width by a factor of 2.5. The consequence is a three-zone decision map rather than a single threshold: differences at or below 1.06 points are non-significant under every compatible table, differences at or above 6.03 points are significant under every compatible table, and in between the published numbers cannot decide. We then show the omission is repairable at negligible cost. One additional integer, the number of units both arms classify correctly, identifies the joint table exactly and restores standard paired inference. For a panel of readers the pairwise dependences must arise from one joint distribution, a constraint that binds once three readers are present; publishing each reader's joint-correct count against a single reference reader cannot widen and may tighten every pairwise bound, and in a 7-arm experiment reduced them by a median of 37% even for pairs excluding that reference. Where the integer was never published we give DentalPair-Cert, an interval with finite-sample coverage uniformly over every admissible within-unit AI-dentist dependence under the independent-sampling-unit model, certified in both the nuisance maximization and the inversion. Across 4,200,000 simulated comparisons an independence analysis falls to 74.5% coverage with 12.2% type-I error; in a purposive sample of 9 recent comparative studies, 1 reported a paired test on discordant units.
Tugrul, M.; Kara, M.
Show abstract
Radiation-induced DNA double-strand breaks (DSBs) drive cellular mortality, mutagenesis, and severe evolutionary bottlenecks. While classical phenomenological models, such as the Linear-Quadratic (LQ) framework, reliably predict macroscopic population survival, they obscure the intrinsic single-cell stochasticity that governs critical rare events like tumor recurrence or the emergence of radioresistant persisters. To bridge this divide, we develop a mathematically exact stochastic differential equation (SDE) framework that models continuous DSB induction and repair as a Feller square-root process. By deriving exact closed-form expressions for the foci moments, we establish a highly efficient Maximum Likelihood Estimation (MLE) pipeline that circumvents computationally exhaustive Monte Carlo simulations, allowing the direct extraction of deterministic repair velocities and intrinsic molecular noise from empirical single-cell $\gamma$-H2AX data. Integrating this kinetic model with a cumulative damage hazard via the Feynman-Kac formalism, our framework seamlessly recovers the classic macroscopic LQ survival topology from microscopic first principles. Furthermore, systematic sensitivity analysis uncovers a fundamental evolutionary duality: while initial physical damage operates additively, ultimate cellular fate is driven by a nonlinear survival response governed by the trade-off between the damage hazard rate and intrinsic molecular noise strength. Crucially, we demonstrate that this molecular noise inherently enhances population survival. Governed by Jensen's inequality, stochastic variance acts as a non-genetic bet-hedging mechanism that buffers the population by favoring cells with transiently low damage loads. Ultimately, this exact stochastic framework bridges microscopic biophysics and macroscopic demographics, offering deep mechanistic insights into the evolutionary roots of radioresistance.
Phan, T.; Pagane, N.; Kreig, J. A. F.; Marc, A.; Locke, M.; Peluso, M. J.; Sandel, D. A.; Deitchman, A. N.; Rutishauser, R. L.; Deeks, S. G.; Ke, R.; Ribeiro, R. M.; Perelson, A. S.
Show abstract
A key goal in HIV-1 cure research is to understand why some individuals control viral rebound after stopping antiretroviral therapy (ART). Recent human studies have identified responding CD8+ T cells expressing Ki-67 and the transcription factor TCF-1 as correlates of post-treatment control, but the mechanistic basis of this association remains unclear. Using the theoretical framework of Conway and Perelson, we fit mechanistic within-host models to viral load and CD8+ T cell data from 9 individuals in a combination immunotherapy trial following ART interruption. Although Ki-67 and TCF-1 measurements were not used for fitting, the inferred effector cell expansion sensitivity, i.e., the responsiveness of effector expansion to low antigen levels, shows a strong linear relationship with Ki-67 and TCF-1 levels at rebound (Pearsons r {approx} 0.8). Building on this, we show analytically that the post-rebound viral load set point is inversely proportional to the effector cell expansion sensitivity, and thus strongly correlates with cycling (Ki-67+) CD8+ T cells (r {approx} -0.8) at rebound, and a subset that expresses TCF-1 (r {approx} -0.9). In effect, individuals with a larger proportion of CD8+ T cells responding to viral rebound, and a greater representation of TCF-1 expressing cells within the responding subset, achieve markedly lower viral set points through a higher effector cell expansion sensitivity. This mechanism is consistent with prior modeling in a non-intervention ATI setting, suggesting it may generalize across more rebound contexts. Our results provide a mechanistic explanation why both Ki-67+ responding CD8+ T cells and their TCF-1-expressing subset predict post-treatment control, linking clinical correlation to its underlying cause and highlighting Ki-67 and TCF-1 as potential early biomarkers of HIV immunotherapy success.
Frost, H. R.
Show abstract
We describe LRSPAT (low-rank spatial toolkit), a fast and memory-efficient framework for approximating measures of spatial association for high-dimensional data. While LRSPAT can be applied to any multivariate spatial dataset, development was motivated by the computational challenge of identifying spatially variable genes in high-resolution spatial transcriptomics (ST) data generated by technologies such as 10x Visium HD, Xenium and Atera. LRSPAT leverages a truncated SVD of the expression data and a thresholded spatial weights matrix to perform reduced-rank reconstruction of spatial statistics in the quadratic form family, including global and local versions of Moran's I, Geary's C, and Getis-Ord G. A regularization approach is leveraged to account for the inflated null distribution of spatial statistics computed on latent variables. By performing key operations on the low-dimensional embeddings, LRSPAT is orders of magnitude faster than standard implementations with significantly lower memory requirements. Because the low-rank approach denoises and desparsifies ST data, LRSPAT is also more accurate than standard techniques at identifying genes with true spatial expression patterns. The dramatic improvements in execution time and memory consumption enable the genome-wide analysis of spatially variable genes (SVGs) and exploration of the full range of hyperparameters including spatial scale, distance metric, and embedding rank. This preprint outlines the background and mathematical details of the approach with limited preliminary results and a short conclusion.
Chen, Y.; Yi, H.; Rao, S.; Weber, A.; Hassmiller-Lich, K.; Sylvia, S.
Show abstract
Inappropriate antibiotic use presents a major global health challenge, particularly in low-resource settings where access to quality care is limited but antibiotics remain relatively unrestricted. This study estimates the causal effect of frontline primary care quality on inappropriate community antibiotic use, combining detailed community-based data from approximately 100 rural villages in rural China with an instrumental variable (IV) approach embedded within a double/debiased machine learning (DML) framework. We linked objective measures of village doctor clinical practice quality, measured through unannounced standardized patient visits, to household-level antibiotic use data collected from the same villages. To identify the causal effect, we constructed multiple candidate instruments from extensive provider characteristics and used an ensemble of machine learning algorithms within a flexible DML-IV framework to approximate an optimal instrument, addressing a many-weak-instruments problem. We found that improving village provider clinical practice quality reduced both antibiotic receipt during healthcare encounters for common diseases and household antibiotic storage for future self-medication. Our findings suggest that strengthening frontline primary care quality can meaningfully reduce inappropriate community antibiotic use without restricting access to essential treatment. More broadly, this study illustrates how causal machine learning can strengthen conventional causal estimation in complex observational settings in global health economics research.
Koshute, P.; Fagan, W. F.
Show abstract
Ecologists remotely track movement steps of animals (e.g., via global positioning systems) and use step selection functions to study the effect of environmental factors upon their movement decisions. Constructing such functions requires pairing each observed step with some number of unobserved but feasible comparison steps. Larger numbers of comparison steps generally yield better estimates but also incur potentially challenging computational demands. Thus, it is important to determine an appropriate number of comparison steps. No established guidance exists for this decision. Here, we use simulated tracks to assess how many comparison steps are needed, fitting each set of steps to a conditional logistic regression model. We monitor errors in estimated effects for several classes of tracks, identifying the number of comparison steps for which mean relative absolute error in estimated effects is consistently low. By this criterion, 32 comparison steps per observed step are needed for our primary class of simulated tracks. Tracks in more homogeneous landscapes, tracks with shorter mean step lengths, or shorter tracks generally require more comparison steps (ranging from 64 to 128 per observed step) to achieve the same level of accuracy. Longer tracks generally require fewer comparison steps (16 per observed step). These results clearly demonstrate that the number of comparison steps influences how well step selection functions estimate covariate effects and provides initial direction in a research area that currently lacks quantitative guidance. Movement ecologists should take care when selecting the number of comparison steps paired with each observed step because those decisions matter.
Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.
Show abstract
Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.
Show abstract
Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.
Zhuang, H.; Zakama, A.; Heller, K.; Faulkner, S.; Gollub, B.; Young-Lin, N.; Chen, I. Y.; Asiedu, M.
Show abstract
In this work, we demonstrate the unprecedented value of NIH's "All of Us Research Program" (AoURP) dataset in studying maternal morbidity and building predictive machine learning (ML) models across heterogeneous populations in the United States. We developed robust and data-driven preprocessing pipelines to curate a longitudinal, multi-site, multimodal, and demographically diverse pregnancy dataset (20,253 subjects; 27,525 pregnancy episodes) from AoURP data, using electronic health records (EHR) (Conditions, Labs, Measurements) and survey responses (Social Determinant of Health (SDoH)), focusing on 7 crucial maternal health adverse outcomes. After characterizing data quality, missingness, and heterogeneity, we performed statistical correlation analysis to identify risk factors. We subsequently developed XGBoost and sequential LSTM models to predict the adverse outcomes, reaching state-of-the-art performance for multiple outcomes. We conducted model interpretability post-hoc analysis to understand success points and fairness analysis to evaluate implications for socio-economic disparities. Four practicing physicians reviewed the set of statistically significant and ML model identified features to assess their clinical validity and novelty. Most features identified through either statistical correlations or ML feature importance analysis aligned with known clinical risk factors. Several features were identified that the ML models used but that are not currently used in clinical practice and may merit further clinical investigation. Fairness analysis revealed certain associations with SDoH and age highlight areas that warrant continued monitoring. Overall, we demonstrate that meaningful populational level patterns can be extracted, and high-performing machine learning models can be trained on this longitudinal, diverse, multi-site dataset. Important risk features, particularly novel ones identified, if validated, could inform new strategies for maternal care or enable development and validation of outcome-specific, clinically deployable ML models.
Bandini, V.; Whitaker, L. H.; Vincent, K.; Salmeri, N.; Mawson, R.; Vercellini, P.; Horne, A. W.
Show abstract
Background: Endometriosis is a chronic pain condition in which hormonal therapies form the cornerstone of long-term management. Treatment tolerability is critical for adherence and therapeutic success, but most comparative studies and reviews have focused on their ability to reduce menstrual pain, while their impact on non-menstrual pelvic pain (NMPP), bleeding patterns, adverse events (AEs), treatment discontinuation and quality of life (QoL) remain poorly characterised. This systematic review and meta-analysis evaluate these outcomes across currently available hormonal therapies, providing practical evidence for clinical decision-making. Methods: PubMed/MEDLINE, Scopus, and Embase were searched up to November 2025 for randomised controlled trials comparing at least two active first- or second-line hormonal treatments for endometriosis. Studies without confirmed endometriosis, treatment duration less than three months and comparing therapies to placebo only were excluded. Data were extracted by two reviewers from reports. Pain outcomes were pooled as mean differences (MD, 95% CI), with bleeding patterns, AEs, and discontinuations as proportions. Analyses were performed in R. PROSPERO: CRD420251137785. Findings: Of 1892 records screened, 48 trials (5583 women) met our inclusion criteria. Overall pelvic pain (0-10 scale) was significantly reduced across all treatment categories (p<0.001): combined oral contraceptives (COCs) (MD 3.17), oral and long-acting progestogens (MD 3.83; MD 4.29), and GnRH-analogues (MD 3.81). Sensitivity analyses restricted to studies reporting NMPP yielded comparable results. GnRH-agonists showed the most favourable bleeding profile, followed by continuous COCs. However, all regimens reported class-specific AEs, including mood changes, nausea, headache, weight gain, and decreased libido (pooled proportions >10%). Overall discontinuation due to AEs was 7.7%, and vaginal bleeding was the leading cause. Heterogeneity across meta-analyses was high. Risk of bias (RoB2) was moderate to high. Interpretation: Given similar reductions in overall pelvic pain across hormonal therapies, treatment decisions should prioritise differences in bleeding profiles, therapy-specific AEs, and QoL. Funding: None.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Gao, C.; Zhang, Y.; He, X.; Yuan, M.; Mou, F.; Zhou, J.; Chen, H.; Wang, H.; Guo, W.; Wei, Y.; Zhang, Z.; Yin, T.; Zhang, C.; Lian, Z.; Zhu, B.; Liu, J.; Zhang, R.; Fu, G.; Onuma, Y.; Wang, D.; Serruys, P. W.; Yi, F.; Tao, L.
Show abstract
BACKGROUND The optimal antiplatelet regimen in patients with acute coronary syndrome (ACS) and multivessel disease undergoing drug-coated balloon (DCB) angioplasty remains unclear. METHODS This was a prespecified subgroup analysis of the REC-CAGEFREE II trial, which was conducted at 41 sites in China and randomized 1948 exclusively DCB-treated participants with ACS to stepwise dual antiplatelet therapy (DAPT) de-escalation or standard DAPT. The primary endpoint was net adverse clinical events (NACE; including all-cause death, stroke, myocardial infarction, revascularization, and BARC type 3 or 5 bleeding) at 12 months. Participants were stratified into multivessel and single-vessel subgroups according to angiographic characteristics. RESULTS Overall, 720/1948 (37.0%) patients had multivessel disease. The multivessel subgroup was associated with a significantly higher risk of NACE compared with the single-vessel subgroup (12.5% versus 6.7%, HR IPTW:1.84, 95%CI:1.35-2.51, P<0.001). No significant interaction was observed between vessel status (multivessel or single-vessel) and treatment allocation with respect to NACE (Pinteraction=0.542). In the multivessel subgroup, NACE occurred in 44/368 (12.1%) and 45/352 (12.9%) in the stepwise de-escalation and standard DAPT groups (HR IPTW:0.95, 95%CI:0.62-1.75, P=0.818), respectively. In the single-vessel subgroup, NACE occurred in 43/607 (7.1%) and 39/621 (6.3%) in the stepwise de-escalation and standard groups (HR IPTW:1.12, 95%CI:0.72-1.70, P=0.611), respectively. For the prespecified hierarchical secondary endpoint, win ratio analyses yielded more wins for stepwise de-escalation in both subgroups. CONCLUSIONS Among patients with ACS undergoing DCB-only angioplasty, those with multivessel disease were associated with a higher risk of NACE than those with single-vessel disease. Stepwise DAPT de-escalation and standard DAPT exhibited similar risk-benefit profiles in both subgroups.
Nkereuwem, E.; Misaghian, S.; Jaganath, D.; Calderon, R. I.; Luiz, J.; Paradkar, M.; Wambi, P.; Castro, R.; Nerurkar, R.; Wang, M.; Wohlstadter, J.; Franke, M. F.; Kampmann, B.; Kinikar, A.; Zar, H. J.; Segal, M.; Kato-Maeda, M.; Collins, J. M.; Swaney, D.; Cattamanchi, A.; Ernst, J. D.; Wobudeya, E.; Sigal, G.; The Combo Study,
Show abstract
Background. Urine-based testing offers a promising non-sputum approach for diagnosing paediatric tuberculosis. However, the currently available lipoarabinomannan (LAM) assay shows limited sensitivity in children and is primarily indicated for those living with HIV. Co-detection of LAM with Mycobacterium tuberculosis (Mtb) proteins in urine could provide complementary pathogen-derived biomarkers that improve diagnostic performance. Methods. We developed an ultrasensitive multiplex electrochemiluminescence (ECL) immunoassay to measure Ag85B, CFP-10, ESAT-6, MPT32, and MPT64 in urine. We determined the analytical limits of detection and evaluated the diagnostic performance of individual proteins and LAM using urine samples from children with Confirmed, Unconfirmed, and Unlikely pulmonary tuberculosis enrolled across five high-burden countries (The Gambia, India, Peru, South Africa, and Uganda). Performance was assessed overall, by HIV and nutritional status, and across biomarker combinations. Findings. Urine samples from 630 children were analysed (median age was 4 years [IQR 2-8]; 44% female, 15% living with HIV, 19% underweight, 24% with Confirmed tuberculosis). The ECL assay achieved femtomolar limits of detection (1.5 to 4.0 fM). The sensitivity and specificity of individual Mtb proteins were 12-33% and 98-100%, respectively. Ag85B had the highest sensitivity (33%, 95% CI 26-41) for Confirmed tuberculosis and was similar to LAM. A four-antigen signature (Ag85B, MPT64, MPT32, LAM) was 50% sensitive (95% CI 42-58) and 94% specific (95% CI 90-96), and was significantly more sensitive than LAM alone, in particular among those without HIV. An additional sixteen (10%) of children with Unconfirmed TB had at least one Mtb protein or LAM detected. Interpretation. Multiple Mtb proteins are detectable in paediatric urine with high specificity, and multi-antigen signatures can augment sensitivity versus LAM alone. These findings demonstrate the potential of multi-antigen urine detection for childhood TB and define analytical targets for the development of future point-of-care diagnostics. Funding. National Institutes of Health.